Papers with video captioning model
Semi-Supervised Learning for Video Captioning (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing video captioning algorithms are heavily dependent on supervised training data. |
| Approach: | They propose to train the video captioning model on labeled and unlabeled data jointly in a semi-supervised learning manner. |
| Outcome: | The proposed model outperforms state-of-the-art semi-supervised learning approaches on VATEX, MSR-VTT and MSVD datasets. |
Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning (D19-1)
Copied to clipboard
| Challenge: | Existing studies on video captioning focus on the association relationships between multiple modalities. |
| Approach: | They propose a video captioning model with high-order cross-modal attention (HOCA) they propose low-rank HOCA which adopts tensor decomposition to reduce the space requirement . |
| Outcome: | The proposed model captures cross-modal interaction of different modalities and reduces space requirement. |